Papers by Junyi Jessy Li
Copied to clipboard
| Challenge: | Existing diversity evaluation focuses primarily on word-level features. |
| Approach: | They propose a method for evaluating diversity over syntactic features to characterize general repetition in large language models. |
| Outcome: | The proposed method shows that models produce templated text in downstream tasks at a higher rate than what is found in human-reference texts. |
Copied to clipboard
| Challenge: | Existing evaluation metrics poorly approximate parser quality, says a new study . questions under discussion is a linguistic framework that views discourse as asking questions and answering them . |
| Approach: | They propose a framework for automatic evaluation of QUD parsing . they use a dataset of fine-grained evaluation of 2,190 QUD questions . |
| Outcome: | The proposed framework shows that satisfying constraints of QUD is still challenging for modern LLMs. |
Copied to clipboard
| Challenge: | Pre-trained language models have shown impressive results when fine-tuned on large summarization datasets. |
| Approach: | They analyze the training dynamics for generation models, focusing on summarization . they find that a propensity to copy the input is learned early in the training process . |
| Outcome: | The proposed model learns at different stages of fine-tuning, the authors show . they show that factual errors are learnt in later stages, but not at high-loss tokens . |
Copied to clipboard
| Challenge: | Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. |
| Approach: | They propose a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs. |
| Outcome: | The proposed framework characterizes and recovers simplification-induced information loss in form of question-and-answer (QA) pairs. |
Copied to clipboard
| Challenge: | Existing work on medical text simplification has focused on monolingual settings . important findings in medicine are typically presented in technical, jargon-laden language . text simulating models can generate viable simplified texts, but there are outstanding challenges . |
| Approach: | They propose a dataset for medical text simplification in four languages . they evaluate fine-tuned and zero-shot models across these languages based on human assessments and analyses . |
| Outcome: | The proposed dataset evaluates models in English, Spanish, French, and Farsi . it shows that the models can generate viable simplified texts, but there are challenges . |
Copied to clipboard
| Challenge: | a new study examines the use of content addition in text simplification when complex concepts need to be explained. |
| Approach: | They present a data-driven study of content addition in text simplification . they analyze 1.3K instances of elaborative simplification in the Newsela corpus . |
| Outcome: | The proposed study shows that contextual specificity can improve elaboration generation performance. |
Copied to clipboard
| Challenge: | Large language models are capable of producing high quality summaries of general domain news articles in few- and zero-shot settings, but it is unclear whether they are similarly capable in more specialized domains such as biomedicine. |
| Approach: | They use GPT-3 to generate single- and multi-document summaries of biomedical articles, given no supervision, using a set of annotations. |
| Outcome: | The proposed model outperforms fully supervised models in generic news summarization, but struggles to synthesize evidence across multiple documents. |
Copied to clipboard
| Challenge: | a non-linear reading order of academic literature is recognized by authors who make explicit connections between non-adjacent passages. |
| Approach: | They propose an enhanced reading experience which generates questions and searches for answer-bearing passages in academic papers to form intra-document connections when answers are found. |
| Outcome: | The proposed interface makes connections between related but non-adjacent passages even if the author did not make them explicit. |
Copied to clipboard
| Challenge: | Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication. |
| Approach: | They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories. |
| Outcome: | The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes. |
Copied to clipboard
| Challenge: | Existing work on automatic prediction of cognitive appraisals has focused on physiological aspects of emotions. |
| Approach: | They present a dataset that assesses 24 appraisal dimensions across 241 Reddit posts . they find that open-source models fail to automatically assess and explain cognitive appraisals . |
| Outcome: | The proposed dataset assesses 24 appraisal dimensions across 241 reddit posts. |
Copied to clipboard
| Challenge: | a novel approach to update comments based on code changes is proposed . a dataset of open-source software projects is used to train and evaluate the model . |
| Approach: | They propose an approach that learns to correlate changes across two distinct language representations to generate a sequence of edits that are applied to the existing comment to reflect the source code modifications. |
| Outcome: | The proposed model outperforms baselines and automatic metrics with respect to making edits. |
Copied to clipboard
| Challenge: | Past work has focused on word frequency-based approaches to improving specificity, such as penalizing responses with only common words. |
| Approach: | They propose to rerank a sequence-to-sequence model to improve the informativeness, reasonableness, and grammatically of responses by using externally-trained classifiers targeting each of these factors. |
| Outcome: | The proposed model improves the informativeness, reasonableness, and grammatically of responses. |
Copied to clipboard
| Challenge: | Existing methods for simplification of medical texts are limited due to jargon and technical content. |
| Approach: | They propose to automate the simplification of medical texts by penalizing decoders for producing "jargon" terms. |
| Outcome: | The proposed method improves on existing heuristics by penalizing the decoder for producing "jargon" terms. |
Copied to clipboard
| Challenge: | Using a new corpus of sentences from Hindi short stories, we analyze the annotations for five different discourse modes argumentative, narrative, descriptive, dialogic and informative. |
| Approach: | They propose to annotate sentences from Hindi short stories for five different discourse modes argumentative, narrative, descriptive, dialogic and informative. |
| Outcome: | The proposed corpus has a high inter-annotator agreement (0.87 k-alpha) and is able to capture the nuances of the embedded discourse structures. |
Copied to clipboard
| Challenge: | Pre-trained systems are able to capture advice better than rule-based systems, but advice identification is challenging. |
| Approach: | They analyze a dataset of advice posts on two reddit forums and annotate whether they contain advice. |
| Outcome: | The proposed models show that pre-trained models capture advice better than rule-based systems, but advice identification is challenging. |
Copied to clipboard
| Challenge: | Existing discourse formalisms require large taxonomies of discourse relations to be accurate. |
| Approach: | They propose a linguistic framework for discourse analysis using questions under discussion . they propose qUD parser that derives a dependency structure of questions over full documents . |
| Outcome: | The proposed model is trained on a large, crowdsourced question-answering dataset. |
Copied to clipboard
| Challenge: | Recent approaches trained supervised models to detect emotions and explain emotion triggers via abstractive summarization, but this can block necessary responses. |
| Approach: | They propose to augment an abstractive dataset with extractive triggers and develop unsupervised models that can jointly detect emotions and summarize their triggers. |
| Outcome: | The proposed model outperforms existing models and is based on a COVID-19 crisis dataset. |
Copied to clipboard
| Challenge: | Recent work explored long-form answers, where answers are free-form texts consisting of multiple sentences. |
| Approach: | They develop an ontology of six sentence-level functional roles for long-form answers . they annotate 3.9k sentences in 640 answer paragraphs and train a strong classifier . |
| Outcome: | The proposed model-generated answers agree less with model-driven answers than human-written answers. |
Copied to clipboard
| Challenge: | Discourse signals are often implicit, leaving it up to the interpreter to draw inferences . current discourse data and frameworks ignore the social aspect, expecting only a single ground truth . elisa f. and her team present a dataset with multiple and subjective interpretations of English conversation . |
| Approach: | They present a first discourse dataset with multiple and subjective interpretations of English conversation . they show disagreements are nuanced and require a deeper understanding of contextual factors . |
| Outcome: | The proposed dataset shows disagreements are nuanced and require deeper understanding of contextual factors. |
Copied to clipboard
| Challenge: | Existing work on bias in NLP only considers negative or pejorative language use. |
| Approach: | They propose a revised framing of bias in terms of intergroup social context and its effects on language output. |
| Outcome: | The proposed framework is based on a model of intergroup relationships in English language tweets. |
Copied to clipboard
| Challenge: | Recent work in NLP has taken advantage of question generation capabilities of LLMs to enhance a wide range of applications. |
| Approach: | They propose a salience predictor for inquisitive questions that is instruction-tuned . they show that highly salient questions are empirically more likely to be answered in the same article . |
| Outcome: | The proposed model is based on linguist-annotated salience scores of 1,766 questions . it shows that answering salient questions improves comprehension of the text . |
Copied to clipboard
| Challenge: | Discourse structure is integral to understanding a text and is useful in many NLP tasks. |
| Approach: | They propose a structured attention mechanism for text classification that derives a tree over a text, akin to an RST discourse tree. |
| Outcome: | The proposed model improves performance on multiple discourse-relevant tasks and datasets and ablation studies show it does little to capture discourse structure. |
Copied to clipboard
| Challenge: | Large Language Models excel at text summarization, but the exact notion of salience remains unclear. |
| Approach: | They propose a framework to derive and investigate information salience in Large Language Models (LLMs) using length-controlled summarization as a behavioral probe into the content selection process. |
| Outcome: | The proposed framework derives a proxy for how models prioritize information in large language models. |
Copied to clipboard
| Challenge: | a dataset of 112 admissions instructions is used to simplify the language used by higher education institutions to communicate with prospective students. |
| Approach: | They propose to simplify admissions instructions by professionally simplifying them and comparing them to a dataset of 112 admissions documents. |
| Outcome: | The proposed dataset includes 112 admissions instructions from higher education institutions across the US. |
Copied to clipboard
| Challenge: | Recent work shows that natural language context is useful in guiding bug-fixing models, but requires prompting developers to provide this context. |
| Approach: | They propose to use bug report discussions to prompt developers to provide natural language context for bug-fixing models. |
| Outcome: | The proposed approach reduces the need for additional information from developers. |
Copied to clipboard
| Challenge: | Using context + knowledge of discourse connectives to make predictions about discourse connective . |
| Approach: | They present a dataset of 8,880 stimuli that evaluates LMs’ inferences about novel entities in contexts where connectives link the entities to particular attributes. |
| Outcome: | The proposed dataset evaluates LMs’ inferences about new entities in contexts where connectives link the entities to particular attributes. |
Copied to clipboard
| Challenge: | Discourse particles are crucial elements that subtly shape the meaning of text. |
| Approach: | They examine the capacity of linguists to distinguish fine-grained senses of English *just* . they find that they struggle to fully capture more subtle nuances of discourse particles . |
| Outcome: | The study shows that linguists struggle to capture subtle nuances of discourse particles. |
Copied to clipboard
| Challenge: | Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. |
| Approach: | They hypothesize that context plays an important role in accurate human annotation and add uncertainty measures can improve model accuracy and calibration. |
| Outcome: | The proposed model can be better calibrated by adding uncertainty measures to models with better accuracy and calibration. |
Copied to clipboard
| Challenge: | SNaC framework is used to evaluate long summaries, but it fails to identify gaps in coherence . nallapati and colleagues have developed a framework for fine-grained annotations of long summarizations . |
| Approach: | They propose a narrative coherence evaluation framework for fine-grained annotations of long summaries that can be used to evaluate coherent narratives. |
| Outcome: | The proposed framework can support future work in document summarization and coherence evaluation, the authors show . |
Copied to clipboard
| Challenge: | During natural disasters, people often use social media platforms to express contempt or sarcasm . despite being widely researched as an NLP task, sarkasmatic detection has not been explored in a specific context . |
| Approach: | They propose a dataset of 15,000 tweets annotated for intended sarcasm . they propose sarkasmatic detection using pre-trained language models . |
| Outcome: | The proposed model can obtain as much as 0.70 F1 on the dataset. |
Copied to clipboard
| Challenge: | Existing work has sought to identify what triggers or causes a particular emotion, but the relationship between those triggers and the prediction of emotion detection models is little understood. |
| Approach: | They propose a dataset to evaluate the ability of large language models to identify emotion triggers . they compare features considered important for emotion prediction models to those considered less salient . |
| Outcome: | The proposed dataset compares large language models and fine-tuned models on social media posts . it shows that emotion triggers are not considered salient features for emotion prediction models . |
Copied to clipboard
| Challenge: | Existing methods for emotion detection are limited in disaster-centric domains due to distributional shifts. |
| Approach: | They propose to use a Twitter emotion dataset to analyze emotions in natural disasters . they propose to apply classification tasks to discriminate between coarse-grained emotions . |
| Outcome: | The proposed model achieves only 68% accuracy after pre-training with unlabeled Twitter data. |
Copied to clipboard
| Challenge: | Existing models are overwhelmingly accurate when presented with counterfactual medical evidence . prior work explored conflicts between context and LLM parametric knowledge in the general domain . |
| Approach: | They construct a counterfactual medical QA dataset that requires models to answer clinical comparison questions with evidence from randomized controlled trials. |
| Outcome: | The proposed model overemphasizes the latter, and the model overestimates the latter. |
Copied to clipboard
| Challenge: | Existing tools to evaluate long text outputs are lacking in the field of NLP . human rating and error analysis remains a crucial component for any evaluation of long text generation. |
| Approach: | They propose a web-based toolkit to collect fine-grained error annotations for long texts . they use a taxonomy to identify errors and assign them to text spans . |
| Outcome: | The proposed tool can be used to evaluate the coherence of long generated summaries. |
Copied to clipboard
| Challenge: | Recent research has made great strides towards understanding the ideological bias (i.e., stance) of news media along the left-right spectrum. |
| Approach: | They propose a novel approach for the study of ideology based on its left or right positions on the issue being discussed. |
| Outcome: | The proposed method allows for the quantitative and temporal measurement and analysis of polarization as a multidimensional ideological distance. |
Copied to clipboard
| Challenge: | In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day. |
| Approach: | They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured. |
| Outcome: | The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials. |
Copied to clipboard
| Challenge: | Software bugs in open-source projects are reported through issue tracking systems like GitHub Issues. |
| Approach: | They propose a method for generating a natural language description of a bug by synthesizing relevant content within the discussion. |
| Outcome: | The proposed system generates a natural language description of the solution by synthesizing relevant content within the discussion. |
Copied to clipboard
| Challenge: | 3,000 English tweets labeled with emotions are used to predict emotions during crises . authors propose semi-supervised learning to bridge this gap . |
| Approach: | They propose to use a dataset of 3,000 English tweets labeled with emotions . they propose semi-supervised learning to bridge this gap by analyzing unlabeled data . |
| Outcome: | The proposed model can be used to predict emotions in the context of COVID-19 . the proposed model performs better than other models using unlabeled data . |
Copied to clipboard
| Challenge: | We model intergroup bias as a tagging task on English sports comments from forums dedicated to fandom for NFL teams . linguistic descriptions of win probability are used for large-scale analysis of intergroup variation . |
| Approach: | They propose to model intergroup bias as a tagging task on NFL fan comments . they use linguistic models to model the bias and use them to generate large-scale annotations . |
| Outcome: | The proposed model can reveal unobserved variations in the form of referents across win probabilities. |
Copied to clipboard
| Challenge: | Existing evaluation methodologies for code summarization tasks do not consider timestamps of code and comments. |
| Approach: | They propose a time-segmented evaluation methodology for code summarization that considers timestamps of code and comments during evaluation. |
| Outcome: | The proposed evaluation methodology compares with other evaluation methodologies that have been widely used. |
Copied to clipboard
| Challenge: | Existing systems for text comprehension are inadequate for more holistic comprehension of a discourse. |
| Approach: | They propose a new paradigm that captures both discourse and semantic links between sentences in the form of free-form, open-ended questions. |
| Outcome: | The proposed model captures discourse and semantic links between sentences in the form of free-form, open-ended questions. |
Copied to clipboard
| Challenge: | Existing work to evaluate LLMs' alignment with human values and opinions has a key shortcoming. |
| Approach: | They propose to add supervision to LLMs to improve alignment with diverse populations . they find that supervision improves alignment across public health, public opinion, values and beliefs . |
| Outcome: | The proposed method improves the alignment of LLMs with diverse populations on subjective questions. |
Copied to clipboard
| Challenge: | Evidence-based medicine connects to every individual, yet the nature of it is highly technical . e-fact-checking systems that connect to medical decisions are largely unused . we examine how clinical experts verify real claims from social media . |
| Approach: | They propose that fact-checking should be approached as an interactive communication problem . they argue that social media and AI have made medical knowledge accessible . |
| Outcome: | The proposed method is based on the work of a clinical expert on social media . it reveals that the method is difficult to connect claims to clinical trials . |
Copied to clipboard
| Challenge: | FactPICO is a factuality benchmark for plain language summarization of medical texts describing randomized controlled trials . existing metrics for factual summarizing medical evidence are poorly correlated with expert judgments on the instance level. |
| Approach: | They propose a factuality benchmark for plain language summarization of medical texts . they assess factuality of critical elements of RCTs in those summaries . |
| Outcome: | The proposed benchmark assesses the factuality of medical summaries using LLMs . the summary summators are based on 345 plain language summaires with fine-grained evaluation . |
Copied to clipboard
| Challenge: | Recent work has explored the capability of large language models to identify and correct errors in LLM-generated responses. |
| Approach: | They propose to combine refinement with feedback into three distinct competencies . step 1: Detect, Critique, Refine gives a fine-grained feedback about errors . |
| Outcome: | The proposed method outperforms existing refinement approaches and models not fine-tuned for factuality critiquing. |
Copied to clipboard
| Challenge: | Automated simplification models aim to make input texts more readable without altering their meaning. |
| Approach: | They propose a taxonomy of errors that are used to analyze simplification models . they propose to use simplification methods to make input texts more readable . |
| Outcome: | The proposed models introduce errors that are not captured by existing evaluation metrics. |
Copied to clipboard
| Challenge: | Vulgarity is a common linguistic expression and is used to perform several linguistic functions. |
| Approach: | They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance. |
| Outcome: | The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings. |
Copied to clipboard
| Challenge: | Existing data-driven questions generate questions that fill gaps in knowledge . a dataset of 19K questions is used to generate meaningful questions . |
| Approach: | They propose a dataset of 19K questions that are elicited while a person is reading a document. |
| Outcome: | The proposed model generates reasonable questions, but the task is challenging. |
Copied to clipboard
| Challenge: | Large-scale crises such as the COVID-19 pandemic cause emotional turmoil worldwide. |
| Approach: | They propose a method to jointly detect emotions and summarize emotion triggers in social media posts related to COVID-19. |
| Outcome: | The proposed method can detect emotions and summarize emotions in long social media posts. |
Copied to clipboard
| Challenge: | Current studies of bias in NLP rely on identifying (unwanted or negative) bias towards a specific demographic group, but this is not always practical. |
| Approach: | They extrapolate a notion of bias from social science literature to predict interpersonal group relationship (IGR) using interpersonal emotions as an anchor. |
| Outcome: | The proposed model predicts the interpersonal group relationship (IGR) using interpersonal emotions as an anchor. |
Copied to clipboard
| Challenge: | Automated text simplification is often thought of as a monolingual translation task . this view fails to account for elaborative simplification, where new information is added into the simplified text. |
| Approach: | They propose to view elaborative simplification through the lens of the Question Under Discussion framework . they propose to model 1.3K elongations accompanied by implicit QUDs to investigate what writers elaborate upon . |
| Outcome: | The proposed framework provides a robust way to investigate what writers elaborate upon, how they elaborate, and how elaborations fit into the discourse context. |
Copied to clipboard
| Challenge: | Neural models for NLP have yielded significant gains in predictive accuracy across tasks. |
| Approach: | They propose a white-box NLP classification architecture based on prototype networks . they propose an interleaved training algorithm that faithfully explains model decisions . |
| Outcome: | The proposed model matches BART-large and exceeds BERTlarge on propaganda detection tasks. |
Copied to clipboard
| Challenge: | a new study examines the use of labeled and unlabeled corpora in political science research . large corporata often contain documents of a certain subject or type, but they are often unlabed . a recent study found that labeles with pertinent documents stem from a single source . |
| Approach: | They propose an unsupervised domain adaptation framework that uses a text classification model and time-aware training to ensure it works well with diachronic corpora. |
| Outcome: | The proposed framework outperforms benchmarks on an expert-annotated dataset and is more stable and learns better representations. |